Papers by Josef Van Genabith
CLaS-Bench: A Cross-Lingual Alignment and Steering Benchmark (2026.findings-acl)
Copied to clipboard
Daniil Gurgurov, Yusser Al Ghussin, Tanja Baeumel, Cheng-Ting Chou, Patrick Schramowski, Marius Mosbach, Josef Van Genabith, Simon Ostermann
| Challenge: | Understanding and controlling behavior of large language models (LLMs) is an important topic in multilingual NLP. |
| Approach: | They propose a lightweight parallel-question benchmark for evaluating language-forcing behavior in large language models across 32 languages. |
| Outcome: | The proposed benchmark measures language steering in 32 languages across 32 languages. |
Why Does Reinforcement Learning Generalize? A Feature-Level Mechanistic Study of Post-Training in Large Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Reinforcement learning (RL)-based post-training often improves the reasoning performance of large language models beyond the training domain, while supervised fine-tuning (SFT) frequently leads to general capabilities forgetting. |
| Approach: | They propose a feature-level mechanistic analysis methodology to probe RL generalization using a controlled experimental setup. |
| Outcome: | The proposed method identifies a compact, task-agnostic set of features that directly mediate generalization across diverse tasks. |
DualFact+: A Multimodal Fact Verification Framework for Procedural Video Captioning (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing evaluation metrics fail to evaluate factual correctness in procedural video captions . Existing metrics rely on lexical overlap or holistic semantic similarity, but miss role-specific omissions resulting in hallucinations . |
| Approach: | They propose a role-aware, fact-level evaluation framework that distinguishes conceptual facts from contextual facts. |
| Outcome: | Experiments show that state-of-the-art captioning models produce fluent but incomplete descriptions with systematic errors. |